sync: merge upstream/main (173 commits) and arm an upstream drift alert - #18
Merged
Merged
Conversation
…kunchenguid#4424) * fix(pr-merge): treat plan-gated 403 on branch rules as no merge queue (kunchenguid#42) * fix(pr-merge): read a plan-gated 403 on branch rules as no merge queue github_read_queue_method left status=unreadable for every failed rules read, including a 403 whose body is GitHub's own "Upgrade to GitHub Pro or make this repository public" message. A repository whose plan cannot expose branch rules cannot have a merge_queue rule either, so that specific 403 now resolves to status=none instead of unreadable - unblocking the away-merge grant on private repos without GitHub Pro. Any other failure (auth, rate limit, network, 404, unrelated 403) still reads as unreadable. * no-mistakes(document): Update stale away-merge queue-grant comment for plan-gated 403 --------- Co-authored-by: NewAiCoder <claude@theinbtw.com> * no-mistakes(review): Fix misleading away-queue-grant comment in fm-pr-merge and its test * no-mistakes(document): Update architecture.md for plan-gated-403 merge queue exception --------- Co-authored-by: NewAiCoder <claude@theinbtw.com>
…unchenguid#4246) * fix(tests): select readers of a changed top-level test fixture bin/fm-test-run.sh --changed recognised shared test helpers by an explicit list, tests/lib.sh|tests/*-helpers.sh|tests/fixtures.sh. A top-level tests/*-fixture.sh matched none of those, fell through to the tests/* catch-all, and was marked unmapped, so selection aborted with "no changed-test mapping for source path" and the run selected nothing at all. tests/herdr-client-pair-fixture.sh and tests/remote-herdr-fixture.sh are real shared fixtures with real consumers, so any branch touching one of them left a validation pipeline driving --changed with a hard abort rather than a narrowed selection. Extend the helper arm to tests/*-fixture.sh rather than routing it through the tests/fixtures/*/* arm. Both arms resolve consumers with the same reference scan, and that scan is what selects the right suites here: it finds exactly the tests that read the fixture. The fixtures/ arm adds only a directory-keying step, which has nothing to key on for a top-level file, so the helper arm is the same behaviour with no extra machinery. A tests/ path nothing reads still reaches the catch-all and still refuses loudly. Refs kunchenguid#4100 * no-mistakes(test): order nested fixtures arm before top-level fixture glob * no-mistakes(document): document tests/ shared-file mapping contract and arm order * no-mistakes(review): drop vacuous test phase, correct header claim, restore comment
… asked, not declined (kunchenguid#4387) * fix(bin): read Claude Code's default external-imports flags as never asked, not declined (kunchenguid#4378) fm-claude-trust.sh refused the whole trust registration whenever the project-root entry carried hasClaudeMdExternalIncludesApproved === false, on the premise that Claude Code writes that value only on an explicit "No, disable". Claude Code's default project entry carries Approved and WarningShown both false before the dialog is ever shown, so every such project refused every spawn. Only Approved === false with WarningShown === true — the pair the dialog writes on a decline — now counts as a decline. false/false behaves like an absent flag: trust is registered and no import consent is manufactured. New case test_project_root_entry_default_import_flags_are_not_a_decline fails on b182d0f with the refusal and passes with the fix; tests/fm-claude-trust.test.sh 31/31, bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 and actionlint 1.7.12. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * no-mistakes(review): Correct harness doc's external-imports decline predicate --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…chenguid#4445) * fix(brief): keep operator address out of composed intent Teach raw-word authoring for intent sections and mid-task relays, with a neutral [captain] provenance marker for legacy mixed tasks. Keep headings and contract prose outside the serialized intent body. The legacy selector already excluded the old speaker labels from its output; preserve that read compatibility. The reproduced leak comes from adding labels inside a modern intent body, not from the legacy selector. Do not scrub actual request content. Add exact serialized-input and generated-contract regressions, retaining refusal of unmarked legacy tasks and coverage of scout promotion. Fixes kunchenguid#3882 * no-mistakes(review): Refuse operator-address lines in Captain's intent body * no-mistakes(document): Document operator-address refusal in intent contract comments
…as a proven empty composer (kunchenguid#4455) * fix(composer): accept Grok title overhang * no-mistakes(review): summary: named Grok overhang constant, doc caveat, restored tmux typed-title coverage
…ailure (kunchenguid#4474) * fix(bin): recover Claude auto-arm after timeout * no-mistakes(document): Add host-timeout signal coverage to autoarm test-coverage list
* fix(spawn): establish Claude task channel authority * no-mistakes(document): Document Claude task-worker control-channel trust in harness-adapters reference
…or pending text (kunchenguid#4458) * fix: guard relaunch exit against pending input * no-mistakes(review): Verifying test run in progress * no-mistakes(document): docs(agent-control): document exit's composer-empty fail-safe guard * no-mistakes(ci): fixed 2 tests broken by approved do_exit fail-safe change (empty-only composer gate). herdr-smoke test's sleep-stand-in never renders a real composer -> updated assertion to expect "not proven empty" refusal instead of stale "did not stop" msg. secondmate-restart fake tmux capture-pane returned bare '> ' glyph (never valid empty proof) -> changed to bordered empty box matching fm-control-relaunch fixture. all 4 related suites pass locally now
…unchenguid#4460) * fix: reconcile diverged secondmate updates * no-mistakes(document): Fix stale fm-update.sh/fm-ff-lib.sh purpose lines in docs/scripts.md * no-mistakes(document): docs: reflect secondmate divergence reconcile in README/SKILL.md
…d#4497) * fix(dispatch): support Codex Luna max effort * no-mistakes(review): use portable CODEX_HOME path in codex effort reference
kunchenguid#4498) * feat(calm): render smooth Unicode swell * feat(calm): make sails asymmetric * feat(calm): use quarter sail glyph * no-mistakes(review): docs: sync calm feasibility sprite passage with approved renderer * no-mistakes(document): docs: sync calm wave phase doc comment * no-mistakes(ci): CI の Lint 失敗は tests/fm-calm-pi-extension.test.sh の test_interactive_terminal_e2e 関数で `boat_narrow_sails` が local 宣言に残っていたことによる ShellCheck SC2034 でした。関数内での参照を確認したところ、狭幅端末の検査は boat_narrow_previous / boat_narrow_direction / boat_narrow_reversed に移行済みで、boat_narrow_sails は代入も参照も一切ありませんでした。そのため local 宣言からこの 1 語のみを削除しました(3315 行目)。Calm の描画実装、他のテストアサーション、ドキュメントは変更していません。検証: bin/fm-lint.sh(ローカル変更ファイルモード)exit 0、CI 相当の `shellcheck --norc --external-sources tests/fm-calm-pi-extension.test.sh` exit 0(SC2034 解消)、`bash -n` 構文チェック通過、actionlint 1.7.12 でワークフロー 3 件 valid。
kunchenguid#4491) * fix: supersede scout delivery brief on promotion * fix: preserve ship safety contract after promotion * no-mistakes(document): Document fm-promote.sh now supersedes brief.md on relaunch
…d stop cleanup dropping accents from a held body (kunchenguid#4471) * fix(bin): let captain holds work on hosts with an older JSON::PP Holding a task for the captain, and the cleanup that keeps a captain-held row open, both fail outright on any host whose JSON::PP defaults allow_nonref off - 2.27202 on a Linux desk is one. Both read a task's body back with `decode_json`, but tasks-axi shows a scalar field as a JSON-encoded bare string, and an older library rejects that whole value with "must be object or array". The consequence is fleet-wide on such a host, not one broken command: a worker there cannot formally record a decision for the captain at all. It can only mention the decision in passing in a status line, where it can be missed - which is how a real decision goes unrecorded. The hold reports that the task lost its hold-set stamp; the cleanup cannot return the row to Queued. Both call sites now ask for allow_nonref explicitly rather than inheriting whatever the installed library defaults to. The second one is worth naming: its `/\A"/` guard reads as deliberate, but a leading quote is exactly the bare-string case that fails, so the guard selects for the failing input rather than protecting against it. The regression case forces the older default back off for every perl the commands spawn, then drives both paths - holding a task that carries a body, and tearing down a captain-held row whose deliverable must still be appended. It also probes that the simulation genuinely rejects a bare scalar, so the case cannot pass vacuously on a lenient host. Each half was verified failing on its own unfixed call site with that site's real error message. Suites: fm-captain-hold-lifecycle 51 cases, fm-backlog-atomicity 99 cases, 0 failures. Verification limit: the mechanism is reproduced and tested, but neither fix is verified against a real JSON::PP 2.27202 host, because none is in the loop. This laptop runs 4.06, where the bug does not manifest. `bin/fm-procevent-lavish.sh:471` was checked and left alone - it matches a brace-delimited object before decoding, so allow_nonref never applies. * fix(bin): stop cleanup silently dropping accented characters from a held body Cleanup rewrites a captain-held row's body to append the finished work's deliverable, and the decoder it reads that body with printed decoded characters to a stream with no `:raw` layer. A character at or below U+00FF then came out as one latin-1 byte instead of two UTF-8 ones, so a body reading "café" lost the accent. `fm_backlog_retain` writes that body straight back through `--body-file`, and nothing reported an error - the character was simply gone from a row still waiting on the captain. The decoder now writes bytes, the same `binmode STDOUT, ":raw"` plus `utf8::encode` that the sibling decoder in `bin/fm-captain-hold.sh` already used. Review of the parent commit found this on one of the lines that commit already changed. It predates that change. The test asserts bytes rather than decoded strings, because comparing strings cannot tell latin-1 from UTF-8. It uses two separate rows on purpose: any character above U+00FF makes perl print the whole string as UTF-8, so one body carrying both an accent and an em dash passes even unfixed and proves nothing. Verified failing before the fix on the accented row, passing after. Suites: fm-captain-hold-lifecycle 52 cases, fm-backlog-atomicity 99 cases, 0 failures. * no-mistakes(document): record body-decode regression proofs in captain-hold lifecycle doc * no-mistakes(review): drop whole-file UTF-8 check from retained-body test * no-mistakes(review): correct stale JSON::PP fleet-host claim in lifecycle doc * no-mistakes(review): anchor native-reproduction claims per defect in lifecycle doc
…furniture (kunchenguid#4532) * fix(composer): read codex 0.154's idle starfield and status footer as furniture codex-cli 0.154.0 animates a braille "starfield" around its idle composer: on the row above the bold `›` prompt row, on the `›` row behind the SGR-2 dim `Ask Codex to do anything` placeholder, and on the row below it, then draws a bright status footer (`<model> <effort>[ fast] · <path> · <title>`). The cells are truecolor greys on both sides of the ghost luminance ceiling, so the brighter ones survive ghost stripping, and the rows below the glyph carry no structural edge. The shared classifier selected the bare `›` shape, extended its wrap region over the two rows beneath the glyph, read the survivors and the footer as wrapped typed input, and answered `pending`; the steering doorbell defers on exactly that verdict, so no doorbell ever reached an idle codex 0.154 pane. bin/fm-composer-lib.sh now recognises that furniture by shape, declared once next to the idle placeholders and reached from the two wrap-region boundary points: - a row whose non-whitespace content is entirely braille cells (U+2800..U+28FF, detected byte-exactly under LC_ALL=C) is furniture: it never counts as wrapped typed content and bounds a bare composer's wrap region; braille behind the glyph row's content is stripped before the emptiness decision when nothing else follows the glyph; a row mixing braille with other text stays typed content; - the codex status footer bounds the wrap region exactly as omp's status row does, anchored on the effort token, a spaced middle dot, and a `~` or `/` path cell, so a typed `fix · tests` stays composer input; - `^Ask Codex to do anything$` joins the verified idle-placeholder set; the ghost strip remains what proves that row empty, and the bare-row rule that bright placeholder text is real input is unchanged. Unchanged: the strict blank-row rule, the styled=0 degradation (a plain cmux/orca capture of this screen still reads `unknown`, never `pending`), FM_COMPOSER_GHOST_LUMA_MAX, and every other harness's shape. tests/fm-composer-lib.test.sh carries both live Herdr samples byte-for-byte with the divergence (letters in place of the starfield read `pending`) and the over-stripping negatives; tests/fm-composer-codex-idle-live-e2e.test.sh is the default-on live guard (token-free, skips explicitly without codex or tmux) that launches the installed codex idle and asserts `empty` through both the tmux and the cursorless styled reads, naming codex --version on failure. docs/verification/runtime-backends.md records the dated Herdr evidence: `pending` before, `empty` after, on the captured screen. * no-mistakes(review): drop unreachable codex footer rule and inert placeholder entry --------- Co-authored-by: Todd Billings <todd@usdvcapital.com>
* fix(bin): refuse empty text steers in fm-send A marked secondmate request sent with an empty message delivered only marker and correlation bytes and minted a pending-reply expectation the parent could never see resolved, stalling the fleet with no loud error (kunchenguid#4255). Fail closed on an empty or whitespace-only message on the text path, mirroring the existing --resolve-key refusal. * chore: retain ambient Pi-lens autoformat as its own commit Formatting-only edits produced by ambient Pi-lens autoformat during the msg-loss investigation, kept separate from the behavioural change in c23acba so the fix stays reviewable on its own. AGENTS.md is deliberately excluded: its only autoformat edit stripped the trailing space from the documented FM_OPERATIONAL_PREFIX value, which bin/fm-operational-input.sh:28 defines as "FIRSTMATE_OP: " and line 11 records as permanent compatibility. Documenting that constant without its trailing space makes the doc wrong about the contract, so that one line was restored rather than retained.
…chenguid#4554) On rose-pine-moon the two-color water (cyan crests over blue troughs) read as a pink stripe over aqua, the yellow left sail and mast clashed with the red right sail, and the hull carried a blue interior run. Every water cell is now blue so the swell reads through glyph height alone, and both sail halves, the mast, and the whole hull are one yellow run. Geometry, cadence, animation, direction flip, resize clamping, and the narrow fallback are unchanged. Update the unit and real-TUI color assertions to the new palette and the Calm docs that described the old one.
…chenguid#4270) * fix(watch): stop aging a second mate's active turn from its launch The parent watcher's second-mate wake-loop stall check exempts a mate that is demonstrably inside an active turn, but secondmate_in_active_turn asked busy_turn_over_age first and returned "not in a turn" whenever that said the bound was crossed. busy_turn_over_age ages from state/<task>.turn-ended, falling back to state/<task>.meta. A second mate's turns end in its own home, so the parent never gets a turn-ended mark for it and the fallback ages the mate's last launch. Every mate launched more than BUSY_TURN_MAX_SECS ago was therefore permanently "over age", the busy pane was never consulted, and any turn outstripping FM_SECONDMATE_WAKE_STALL_SECS raised a false wake-loop stall. The gate now bounds the busy exemption by <idle> - how long the queue's drain position has not moved - which is evidence this home actually holds. A busy mate stays exempt while the queue has been frozen for less than BUSY_TURN_MAX_SECS, and a mate stuck busy forever still alarms, so the bound that stops a busy pane from proving liveness forever is kept rather than removed. busy_turn_over_age is untouched; its remaining callers are the ordinary crew busy-pane bound. The regression pins the case that actually broke: a mate whose launch record predates BUSY_TURN_MAX_SECS and which is demonstrably mid-turn must not escalate, while the same mate with its queue frozen past the bound still publishes exactly one notification. The existing coverage only exercised a freshly launched mate, which passes either way. Reaching that alert now costs a pane capture inside the gate, so the three checkpoints in this suite that assert an alert move from a 1s to a 4s bound - the value the neighbouring active-turn cases already use. The bound is a ceiling, not a wait: the checkpoint returns on the first actionable wake. On a loaded machine a 1s bound missed the alert repeatedly; at 4s it did not miss in 20 runs under the same load. * no-mistakes(review): scope the second-mate active-turn regression test's coverage claim * no-mistakes(document): fix stale second-mate active-turn comments in fm-watch
…unchenguid#4278) * feat(bin): add read-only PR blocker and reviewer-discovery commands Two focused, opt-in commands that read GitHub and never write to it. fm-pr-state.sh reports what still blocks one pull request from the author's side: a closed or merged state, draft state, unknown or conflicting mergeability, absent or failing required checks, and a blocking CHANGES_REQUESTED decision explained by each reviewer's latest verdict, marked STALE when it was left at a superseded head. A pull request that only awaits an approval is not reported as blocked, and advisory checks are omitted. Every reading is taken against one exact head; a push that lands mid-read invalidates the whole result rather than mixing two snapshots. fm-pr-reviewers.sh suggests reviewers from the most recent commits to the pull request's exact changed paths, counting each commit once, resolving handles through GitHub's own commit author.login mapping, and excluding the author and Bot accounts. Both stay read-only: no review request, no approval, no merge. Unresolved review-thread state is left unreported because the REST API does not expose it and unattended commands may not use GraphQL. Closes kunchenguid#3731 * no-mistakes(review): accept only PR URLs and stop at terminal state * no-mistakes(review): report unconfirmed required checks; make URL-only guards discriminate * no-mistakes(review): stop attributing readings to unverified heads * no-mistakes(review): narrow readiness contract to checks that have reported * no-mistakes(review): read the pull request once, drop the head guard * no-mistakes(document): scope pr-forge isolation proof to its measured members * no-mistakes(document): record uncovered pr-forge members and their pending proof * docs(isolation-proof): re-prove pr-forge at its full membership tests/fm-pr-state.test.sh and tests/fm-pr-reviewers.test.sh joined the pr-forge family in this branch, and script_allows_concurrency grants four workers by family membership alone, so both ran concurrently on a proof measured before they existed. Re-proved the family at all eight members: two consecutive runs, 0 failures, each begun with the one-minute load average below 6.0 so the result measures isolation rather than contention. A third run taken between them is disclosed rather than recorded, because it started while the previous run's workers were still decaying. The new durations are not comparable with the six-member measurement above them, so they are not presented as evidence about the two new members, and that record's 1.72x four-worker figure is left as a statement about its own run rather than restated as current. * no-mistakes(review): disclose gh error-text coupling at its matching site and tests
…uid#2752) * fix(bin): teach validation-round pauses in briefs * no-mistakes(document): Point classifier comments to authoritative pause examples
…guid#4510) * fix(teardown): refuse a cleanup whose endpoint close failed bin/fm-teardown.sh discarded both the exit status and the stderr of every fm_backend_kill call, so a close that genuinely failed was indistinguishable from one that succeeded. Teardown continued past it, deleted the task's durable records, returned its worktree, and reported the cleanup as completed. The deleted metadata is the only record of which endpoint belongs to the task, so such a close did not merely leave a stray session behind, it stranded one: nothing was left on disk naming it. The adapters could not carry that signal either. Driven against the real code, every backend arm returned 0 for a genuine failure exactly as it did for an already-exited endpoint, so there was nothing for the four call sites to propagate even once they stopped swallowing it. The tmux arm now resolves a close that did not succeed against the window's exact recorded identity, since kill-window fails the same way for a window that is gone and one that is still there. The Orca arm reports a close its missing CLI never attempted. Both stay silent for an endpoint that is already legitimately gone, and the remaining arms are unchanged: their close-command timing cannot be established without the real Zellij, Orca, and cmux binaries, and a gate that refused ordinary cleanup of an already-exited session would be worse than the defect. docs/verification/runtime-backends.md records what each backend can prove. A reported close failure now reaches teardown's existing retain-and-stop refusal before the records naming the endpoint are removed, matching where the Herdr confirmed-gone gates already sit for the same hazard, and the retained records let a rerun finish once the close works. * no-mistakes(review): refuse unreadable tmux close re-read; honor --force override * no-mistakes(review): drop unreachable Orca force arm; prove CLI-absent close * no-mistakes(document): document endpoint-close refusal in its backend and retirement owners * no-mistakes(ci): The two reported failing checks are NOT code defects. Both "CI" (run 34935529184) and "Require no-mistakes" (run 34935529206) returned conclusion=action_required with zero jobs and 0s duration (run_started_at == updated_at), which is this repo's workflow-approval gate holding the run before any job starts. No job executed, so nothing in the diff could have caused them; two unrelated branches (fm/captain-hold-json-nonref, fm/presenter-core-l1) show the identical shape in the same time window. Verified the change locally instead: bin/fm-lint.sh clean, bin/fm-test-run.sh --check-coverage ok, and all suites the diff touches pass (fm-teardown-endpoint-safety 25/25 including the five new endpoint-close cases, fm-backend-orca, fm-backend, fm-backend-tmux-smoke, fm-backend-cmux, fm-backend-zellij, fm-backend-herdr). Separately, I found and fixed a genuinely flaky test that the phase rules require me to make deterministic: tests/fm-tmux-agent-liveness.test.sh intermittently failed "an idle shell pane must classify dead" (verdict ambiguous, comms=[bash sleep]). It is selected by --changed for this diff, so it would run against this PR once CI is approved. Root cause, established by instrumenting the pane's process group: the idle window was created by `new-session` with no command, so it inherited tmux's default-shell, i.e. whoever runs the suite. ps on the pane tty showed `-zsh` -> `bash` -> `sleep`, all sharing pgid==tpgid, i.e. the host operator's shell configuration spawning a periodic helper directly into the pane's FOREGROUND process group, which is the one surface the classifier reads. `sleep` classifies as `other`, so fg_other=1 and the verdict became `ambiguous` instead of `dead` whenever that helper overlapped the 10s poll window. Every other window in the suite runs an explicit command via new_window; the idle case was the only one whose process group the host defined. Fix (smallest root-cause, test-only, 1 line + explanatory comment): create the idle window with an explicit bare `/bin/sh` (`-- /bin/sh`), the same shell the neighbouring background case already execs. Its foreground group is now exactly one process (verified: `/bin/sh` alone), so no host configuration can inject into it. This flake is pre-existing and NOT caused by this PR: an interleaved A/B showed base commit da5e658 failing the identical case (2/6 runs) alongside head (3/7 runs), and the diff only extracted the tmux inventory read into a helper with identical semantics while never touching fm_backend_tmux_foreground_comms. After the fix: 8/8 consecutive passes, with lint and the coverage guard still clean. Change left uncommitted in the working tree
* feat(calm): ship the Claude Code Calm and sailboat mod behind the function-hooks flag Add .claude/mods/firstmate-calm, a Claude Code mod (function-hooks plugin) that brings Calm to Claude Code: the sailboat replaces the stock working row through a Raster repainted on the sprite's own tick, and tool, tool-group, mid-turn narration, and canonically classified operational user rows draw at zero height. /calm is registered by the hooks module itself and toggles the same per-home config/calm preference the Pi extension uses, so one choice applies on either harness; rows redraw retroactively on toggle and stay hidden across claude --continue. The mod loads only while Claude Code's default-off CLAUDE_CODE_ENABLE_FUNCTION_HOOKS flag is on. Nothing sets that flag in any settings file, and the plugin carries no command file, skill, agent, or classic hook, so it is a complete no-op while the flag is off. The trusted project auto-loads it through an .agents/skills symlink, the only path Claude Code scans for project plugins. Extract the working-ship geometry, bounce track, cadences, and freeze/resume state into a harness-neutral sprite core inside the mod (Claude Code refuses hooks-module imports from outside the plugin folder) and have the Pi widget paint that core's frames as standard ANSI, byte for byte as before; the Pi suite stays green. Classify operational rows through a port of bin/fm-operational-input.sh's classify command guarded by a corpus parity test against the shell owner. Tests: portable Node checks (plugin shape, sprite parity with Pi's rendering, Raster packing, policy, classifier parity), the mod's own claude plugin test suites behind a default-on wrapper, and an opt-in live TUI guard proving the flag-off no-op, the moving boat, hidden rows, the persisted toggle, and resume on Claude Code 2.1.272. Docs: record the version-scoped Claude Code evidence and the three bounded gaps in docs/calm-mode-feasibility.md, describe the Claude Code contract in docs/calm.md, and make the shared preference, layout, and contributor notes harness-neutral. * no-mistakes(review): Preserve colliding final replies and strengthen parser parity * no-mistakes(review): Preserve final replies and strengthen canonical parity checks * no-mistakes(review): Require exact function-hooks opt-in before Calm activation * no-mistakes(review): Clarify Calm module loading and activation boundaries * no-mistakes(review): Reset Calm presentation state across session starts * no-mistakes(document): Refresh Calm session lifecycle documentation * feat(calm): paint the Claude Code working ship in Claude's own theme colors The captain picked the "Claude native" palette for the Claude Code mod's Raster: every water cell takes the spinner blue of the active theme family (#93a5ff dark, #5769f7 light) and the whole boat takes the Claude orange of the stock spinner (#d77757), one water color and one boat color. The family follows the `theme` setting's prefix, read at load through $.config.list and re-read on a config.set of that row, with `auto` and custom themes falling back to the dark set. The Pi extension keeps its standard ANSI blue and yellow, byte for byte. Rename the shared sprite's color classes from hue names to `water` and `boat`, since each harness now maps them to its own colors; geometry, motion, cadence, and the activation gate are untouched. Tests cover both palettes' packing and the family rule under Node, and the plugin kit drives every theme value, a theme change mid-session, the Calm-off pass-through, and inertness of the menu read while the flag is off. The docs describe the Claude Code colors and record the guard passing on 2.1.273. * no-mistakes(review): Use light palette for unresolved Claude themes * no-mistakes(document): Refresh Claude Calm verification evidence
…kunchenguid#4586) * fix(watch): honour a declared wait before wedge-escalating a quiet pane wedge_timer_check escalated on elapsed idle time alone. Nothing asked whether the worker had already said why its pane was quiet, so a lane that declared a bounded external wait climbed the escalation ladder for as long as the wait lasted, and past FM_WEDGE_DEMAND_INSPECT_COUNT every repeat carried demand-deep-inspection - which by its own wording forbids re-absorbing on the run-step or pane state, so the supervisor could not use the evidence that was there either. The generated brief promises that declaring `paused:` buys the long recheck cadence instead of a wedge, but the timer was still reachable while that declaration stood: a crew that declares a wait and then has an active run or busy pane attributed to it is handed to the timer as provably-working. The declaration is what the worker said about its own silence, so it now outranks a liveness verdict that only says something is running. The consult runs in the at-threshold branch that was about to escalate, beside the worktree walk already there, and costs one status-line read. Either status-line record defers to the same FM_PAUSE_RESURFACE_SECS recheck the declared-wait absorber already uses, so the wait is still rechecked and cannot rot invisibly. Which verb declared it decides the wording, because the two block on different people: a `paused:` wait is owed by an external dependency and asks the reader to confirm it still holds, while a `captain-held:` transfer is owed by the captain reading the recheck and asks them to answer or release the hold. A hold is not rechecked at all while the away-posture record exists, as on every other captain-held path, and that absorb arms no throttle so the recheck is owed in full on return. A declared clearing time that has already passed stops counting, and a lane that never declared one keeps the identical escalation schedule, reason, count and demand-deep-inspection wording, so detection and its worst-case time are unchanged. The deferral restarts the idle timer rather than cancelling it, so a lane that stops waiting escalates again within one threshold. A lane quiet because its own validation run is parked at a gate awaiting a human decision is deliberately out of scope: reading that state needs a signal carrying who the wait is on and what clears it, rather than one inferred from a parked verdict that also covers gates awaiting the crewmate itself. Tests pin both directions for each case and were each confirmed to fail with the consult removed. * no-mistakes(document): docs: honour declared waits in stale-escalation docs
* fix(bin): derive passed PR state from PR record A completed no-mistakes run with outcome=passed does not prove the associated pull request merged or closed. A parked gate can be approved on other evidence, so the old crew-state label could report an open PR as merged and make teardown look safe when unlanded work still exists. For passed runs, derive the crew-state detail from the run or task PR identity, accept a matching merge-poll retirement receipt as local merged evidence, and otherwise perform a bounded forge read. If the identity is absent or unreadable, report the run as passed with unknown PR state instead of inventing a merged claim. Fixes kunchenguid#4607 * no-mistakes(review): Add bounded GitLab merge-request state reads * no-mistakes(review): Preserve network-free inactive crew-state scans * no-mistakes(document): Document PR record readers in shared library
kunchenguid#4627) * fix: restore published contribution follow-up (Fixes kunchenguid#4469) * fix(review): Fix contribution freshness and merge actor routing * fix(review): Restore issue triage and scope contribution follow-up * fix(test): test: assert one wake per contribution signal * fix(document): Document contribution follow-up * fix: restore truthful terminal delivery evidence * fix(review): Disclose unsupported contributions and deduplicate watcher wakes * fix(review): Preserve unmeasured unsupported contributions across Bearings * fix(review): Deduplicate shared contribution wakes and isolate diagnostics * fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower
…nguid#4658) * fix(bin): make a remote-reply document gap self-clearing and re-attemptable A remote mate's undelivered document raised a keyed `blocked` decision that nothing could ever resolve, and any `data/*.md` substring in any mirrored line was an unconditional fetch instruction. A mate announcing a report it had not written yet therefore manufactured a permanent, factually false blocker, and its own explanation of the false alarm manufactured more. The reader has no permanence vocabulary: a report still being written refuses exactly like a path that will never exist. So an undelivered document is now a durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`, re-attempted on the next delta and on the channel's own quiet poll, and retired with a matching `resolved` line naming the local copy once it arrives. The cursor still advances and no delta stalls on one bad pointer. Only a structured `report=data/....md` pointer now offers a document, so a path merely mentioned in prose - including one under another home's mirror tree, which is provably not that mate's to serve - is never fetched. Offers are deduplicated across the whole delta, the escalation names each missing document once and carries the reader's own reason instead of discarding it, and a strictly increasing notice ordinal keeps a later escalation from being swallowed as duplicate bytes. A mirrored line still lands once whichever pointer form it was first written under. * no-mistakes(review): Require structured pointer token boundaries * no-mistakes(review): Unify boundary-safe pointer extraction and rewriting * fix(bin): identify a mirrored line independently of its delivery state Two defects in the boundary-safe pointer work. The at-most-once check compared only the all-remote and all-local renderings of a line, so it could not recognize a mixed one. A line offering two documents where only the first was deliverable mirrored as local-plus-remote; once the second arrived, a cursor-loss whole-log recapture rendered the same line all-local, matched neither alternate, and mirrored a second time. A line's identity is now the canonical form every boundary-valid pointer would take once delivered, derived by the same parser that does extraction and rewriting, so it no longer depends on which documents happened to be deliverable at the time. The pointer map was passed to awk through the process environment. A delta may carry up to the configured 1 MiB bound, and an expanded map of delivered pointers can exceed the platform's exec argument limit, so awk would fail to start; because no caller checked, the empty result would have been appended as blank lines while the cursor advanced past dropped status content. The map now travels in a file, and every call site checks the exit status and stops the ingest rather than committing a delta it could not render. Both passes now run once per stream instead of twice per line. * no-mistakes(review): Abort ingest when document pointer extraction fails * no-mistakes(review): Exclude structured cross-home pointers from document transfer * fix(bin): fail open on an undeliverable remote document instead of tracking it Narrow the remote-reply document fix to the scope the diagnosis actually requires, as decided after measuring a simpler alternative. A document the reader cannot deliver now fails open. The mate's line is mirrored with its own pointer, the cursor advances, and one unkeyed note carries the reader's reason. A note never enters the open-decision fold, so it cannot stand open the way the original keyed block did - which removes the never-clearing false blocker by construction rather than by resolving it. That makes the durable self-clearing obligation unnecessary, so it goes: the per-mate pending-documents record, its notice ordinal and resolved announcements, and the poll-side retry. Canonical line identity goes too, and with it a way to silently drop a genuine status line; mirroring is back to at-most-once on exact bytes. The cross-home exclusion goes as well: under fail-open a cross-home report= either fails harmlessly or is a nested remote report this mate genuinely holds, which is now relayed again. Kept: fetching only on a structured report= pointer, the boundary-correct parser, the file-based rewrite map, and checked extraction and rewrite exit status. The parser now scans behind a sentinel byte so a rejected candidate can no longer give the text right after it a false leading boundary. The reported incident is covered end to end: a report path announced in prose before it exists raises no decision, and the report still arrives through the ledger publisher's structured offer once written. * no-mistakes(review): Preserve source-line identity across remote reply replays * no-mistakes(document): Document remote reply transfer and replay semantics * no-mistakes(lint): Fix staging truncation lint checks
* Preserve substantive Calm mid-turn text * no-mistakes(review): Distinguish newline-preserved replies from short narration * no-mistakes(document): Document Calm mid-turn preservation boundaries * no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check`
…#4656) * fix(bin): re-record PR poll identity after a volume device renumber (Fixes kunchenguid#4260) A volume remount can renumber the state filesystem's st_dev while every inode and byte stays the same; APFS does this across a reboot. A poll registration records its sidecar and check as device:inode, so every poll armed before the remount failed strict validation and the watcher refused all of them as unauthenticated state checks until each was re-armed by hand. There are two device comparisons. fm_pr_private_file_valid compares a live file's device with the state directory's device read in the same invocation: it refuses a file that is not on the state directory's own filesystem and already survives a renumber, so it is unchanged. The registration's recorded identity versus the live identity (from kunchenguid#556, reused by the kunchenguid#932 retirement receipt) binds the registration to the exact files published in its own transaction; its device part is what breaks. When strict capture fails, the watcher now proves the device is the only difference: every other artifact check passes (template bytes, both hashes, private mode, single link, live device, metadata), both recorded identities name one device, and each recorded inode equals its live inode. Only then, under the task's control lock, does it rewrite the two identity lines, repeating the whole proof and comparing the registration's file identity and bytes just before the rename, and then capture strictly again. A swapped, altered, re-moded, relinked, split-device, or foreign-device artifact still fails a proof and is still refused, and a pending retirement receipt blocks the rewrite. Reproduction: on macOS a poll armed on an APFS disk image that was detached and re-attached behind another image moved st_dev 16777239 -> 16777243 with inodes, bytes, mode, and link count unchanged; the real watcher refused it on main and reports its merge with this change. The portable regression test rewrites a real registration's recorded device and drives the watcher. Not changed here: the status presentation cursor keys rows by its own device:inode identity in bin/fm-classify-lib.sh, a different helper that needs its own fix; a retirement receipt left by a reboot between its publication and removal still names the old device and stays refused; custom check trust binds only a content hash and is unaffected. * fix(review): Serialize PR poll publication writers * fix(review): Bound PR poll publication lock scope
…d#5546) * fix(bin): classify the stdin program of `bash -s` with operands in the arm policy With -s, sh/bash/zsh read the program from stdin even when operands follow; the operands are only positional parameters. The arm policy treated the first operand as a script path, so heredoc and here-string payloads were never classified and a hidden bin/fm-watch.sh execution was allowed. A protected path in the operand position still fails closed as before. Fixes kunchenguid#1489 * no-mistakes(document): Clarify stdin shell operand documentation * no-mistakes(ci): Captain, fixed `shellInvocation` so `bash -- -s` treats `-s` as a script name, and updated R21 to test the exact command. The targeted policy suite, lint, documentation check, and diff check pass. Both hosted workflows show `action_required` before any jobs ran; that external approval state remains unresolved * fix(bin): keep main's handling of words after a leading `--` Revert the pipeline CI-step change that made the first word after a leading `--` always a script. It turned forms that main denies today into allow (for example `bash -- -c 'bin/fm-watch.sh'`), which is outside kunchenguid#1489 and loosens a fail-closed policy. `--` after `-s` still ends option parsing.
…nchenguid#5695) * fix(bin): strip AI co-author trailers from fleet-launched commits Cursor and other non-Claude runtimes append the trailer after the typed message. A per-task commit-msg hook removes it and leaves human co-authors and the author identity untouched. * no-mistakes(review): Export pane hooksPath override and drop generated-with stripping * no-mistakes(ci): This PR caused all three CI failures, and the fix is test-only: 4 test files change, no product code. **Cause.** `fm-spawn.sh` now installs the AI-trailer strip hooks for every spawn, secondmates included. The installer refuses a worktree that is not a git repository, and the PR deliberately keeps that fail-closed rule because real secondmate homes are firstmate clones. Four test fixtures still gave secondmates a plain directory as their home, so each spawn failed with "not a git worktree ... could not install the AI-trailer strip hooks": - serial 5: `tests/fm-backlog-atomicity.test.sh` ("secondmate spawn failed"). - serial 8: `tests/fm-secondmate-harness.test.sh` ("split: no meta written"). - Herdr: `tests/fm-backend-herdr-launcher-workspace-e2e.test.sh` and `tests/fm-backend-herdr-workspace-per-home-e2e.test.sh`. **Rule that must hold.** Every home a test spawns as a secondmate must be a git worktree. I checked the other places in the changed area: the only secondmate spawns in these tests are the ones listed. The earlier rounds already fixed the other fixtures (`fm-secondmate-liveness`, `fm-secondmate-safety`) the same way. **Fix.** - Each of those four secondmate homes now gets the same `.gitignore` plus `git init -q -b main` that the liveness and safety tests already use. - The two Herdr tests clean up with their own plain `rm -rf "$TMP_ROOT"`, not the shared `tests/lib.sh` helper. Because the installer leaves each `state/<id>.git-hooks` directory read-only, that cleanup printed "Permission denied" and left the directories behind. Both cleanups now restore the owner's write bit on every directory before removing (`find ... -exec chmod u+rwx`), which is what `fm_test_remove_tree` in `tests/lib.sh` does. **Verification.** - `tests/fm-secondmate-harness.test.sh` passes. - `tests/fm-backlog-atomicity.test.sh` passes (99 ok, exit 0). - shellcheck is clean on all four files. - I could not run the two real-Herdr tests locally: the Herdr lab on this host refuses to start because it needs exactly one running default session, and I did not change the host's Herdr state to get around that. Instead I checked their two changed steps directly: the installer succeeds on a home set up the new way, and the new cleanup removes the read-only hooks directory completely. Those two tests will only be proven on CI
…id#5683) * fix(bin): treat Pi's dollar-first cost footer as furniture An idle Pi status row opening with $0.000 was read as a dead-shell prompt, so exit and relaunch refused on an empty composer. * test: wait for the draining holder to exec sleep before reading its identity The procevent drain fixture read fm_pid_identity immediately after backgrounding setsid sleep, racing the child's exec chain. Mid-exec the cmdline can read empty, failing the fixture on a loaded CI runner. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…5534) * fix(bin): refuse a merge when a required check never reported fm-pr-merge.sh built its GitHub refusals only from checks present in statusCheckRollup, so a required check that never ran was simply absent and the merge proceeded on the subset that reported, contradicting its own "every required check green" claim. The GitHub verify now reads the base branch's required contexts from the forge itself - the classic branch protection summary on GET repos/{o}/{r}/branches/{b} and the active ruleset rules on GET repos/{o}/{r}/rules/branches/{b} - and refuses when a required context has no entry in the same rollup, at the same head, that the merge is bound to. Absence reads as unknown, never green. The required-set read joins the existing refusal list, so a draft, a red check, and an unreported required check are all reported together. Could not read vs nothing required: both endpoints need only repository read access. The admin-only GET .../branches/{b}/protection endpoint is deliberately not used: it answers a non-admin token with the same 404 an unprotected branch gets (observed live on kunchenguid/firstmate main with this token), which would read a missing permission as "nothing required". Any failed or malformed read of either source (auth, missing fine-grained permission, rate limit, network, 404, unexpected shape) refuses the merge with a line naming the unreadable source. The one exception is GitHub's plan-gated 403 on the rules endpoint ("Upgrade to GitHub Pro or make this repository public"), which already means "this repository has no branch rules" for the merge-queue reader; that check moves into one shared helper and the classic summary still decides for such a repository. Attended waiver: --allow-missing <check-name> is the twin of --allow-red and follows the same design and recording path: once, separate name argument, waives only that exact unreported required check, still requires every other required check reported and every check green, never waives an unreadable required set, refused while the away-posture record exists, and refused on GitLab. Merge-state BLOCKED policy is unchanged. How this differs from the withdrawn kunchenguid#5353 (read from its diff): - kunchenguid#5353 read the admin-only branches/{b}/protection endpoint and treated its 404 as "no required checks", so for any non-admin token the required set silently read as empty; this change reads the read-access branch summary and treats every failure as unreadable. - kunchenguid#5353 ignored rulesets; this change also reads required_status_checks rules from the effective branch rules. - kunchenguid#5353 made separate per-head REST reads of statuses and check-runs capped at per_page=100 with no pagination; this change checks presence in the same statusCheckRollup view the red-check gate already reads at the verified head. - kunchenguid#5353 stopped at the first unreadable read; this change reports it as one refusal among all the others. - kunchenguid#5353 also claimed kunchenguid#5345 (lock stealing) and changed 39 files, most unrelated; this change is kunchenguid#5344 only. Live proof, read-only (a gh wrapper refused every merge and mutating call): - cli/cli#14474 (trunk requires 3 classic build contexts, none ran): refused, naming build (macos-latest), build (ubuntu-latest), build (windows-latest); with --allow-missing "build (macos-latest)" it still refused, naming the other two. - cli/cli#13665 (ran build (ubuntu-24.04-firewall) instead): refused, naming build (ubuntu-latest). - hashicorp/terraform#39262 (ruleset-required checks absent): refused, naming Code Consistency Checks, End-to-end Tests, Race Tests, Unit Tests. - cli/cli#14485 (all required reported and green): verified; the wrapper blocked the merge call and the pull request read back open. Fixes kunchenguid#5344 * fix(review): Preserve required-check producers and aggregate independent read failures * fix(document): Clarify required-check verification and waiver documentation * fix(bin): match an app-bound required commit status by name The producer-identity check resolved an app-bound required context only against check runs, so a required context that the required app reports as a commit status could never match and always read as "has not reported". A commit status carries no app id to compare, so an app-bound requirement that arrives as a status now matches by name, as before producer binding; check runs keep requiring the configured producer app. Live, read-only: hashicorp/terraform#39262 requires license/cla from integration 865473, reported green as a commit status by the CLA app. The previous head refused it as unreported; this head no longer does, while still naming the four required check runs that never ran there. Refs kunchenguid#5344 * fix(document): Clarify accepted commit-status producer verification limitation --------- Co-authored-by: firstmate-oss <firstmate-oss@kunchenguid.local>
…uid#5696) * fix(bin): never offer a persistent secondmate for teardown The return brief's "Landed, cleanup due" scan listed every state/*.meta record carrying a pr= and a merge-notified marker without regard to kind, so a secondmate record holding a relayed child's merged PR put the mate itself up for "bin/fm-teardown.sh <mate>" cleanup. A secondmate is a persistent worker, never landed work. - bin/fm-afk-return.sh: skip kind=secondmate in the landed-cleanup scan. - bin/fm-pr-check.sh: refuse to record pr= or arm a merge watch on a kind=secondmate record before any side effect; a PR reported on its routed status channel belongs to a task in the mate's own home, which arms its own watch. - bin/fm-watch.sh: a merged result from a poll already armed on a secondmate retires the poll silently - no merge outcome, marker, or wake. * no-mistakes(document): Document secondmate merge-watch and return-brief exclusions * no-mistakes(ci): CI failed in an unchanged watcher-shutdown test whose three-second wait was sensitive to runner load. Increased the wait for both state- and home-deletion cases without changing watcher behavior. The full fm-watch-arm suite passed locally; syntax and diff checks passed
…unchenguid#5702) * fix(bin): refuse unknown dash-leading args in public-posting fm-x scripts fm-x-reply.sh collected any unrecognized argument into the positional pool and took the first one as the reply text, so an invocation like "fm-x-reply.sh <id> --followup --final <text>" posted the literal string "--final" to X and silently dropped the real text. Make argument parsing strict in every script that can post publicly: an unknown dash-leading argument, a dash-leading request_id/task id, a dash-leading option value, or a surplus positional now exits 2 with a usage error before any config load, outbox write, or network call. Reply text starting with '-' is still accepted via --text-file or stdin, and --help is honored wherever it appears instead of becoming text (a --help forwarded through fm-x-followup.sh would have counted as a posted follow-up and mutated the link). fm-x-link.sh and the fm-public-followup scripts already refuse unknown arguments; fm-x-poll.sh takes none. * no-mistakes(review): Refuse surplus follow-up text sources; drop post-ID help branches * no-mistakes(document): Clarify reply and follow-up argument usage * no-mistakes(review): Refuse dash-leading --text-file operands in fm-x-reply * no-mistakes(document): Correct follow-up argument parsing comment * no-mistakes(document): Document dismiss argument rejection in script header
…nchenguid#5701) * feat(bin): latch the supervision host after repeated engine errors Rung 3c-1 of the PR 5631 re-cut: the host copies the Pi branch's broken-session policy. Two consecutive engine errors latch the session; every away wake then reaches main with one supervision-host line for a five-minute cooldown, after which one wake probes the engine, and each failed probe doubles the cooldown up to one hour. A reported turn without an engine error clears it. The latch is kept per main session, engine, and model in state/.supervision-host-health, and the engine conversation now uses the same main-session key, which includes the lock holder's process identity so a recycled pid never shares either. Lifted from the validated 5631 tree and adapted to main's away-only host: the attended recovery line and attended cooldown pass-through are left for the attended core, so a recovery is only logged. * no-mistakes(document): Consolidate supervision-host latch documentation
…kunchenguid#5535) * fix(bin): absorb routine second-mate progress while surfacing routed replies Fixes kunchenguid#2959 A kind=secondmate task's status signal was never absorbable, so a healthy mate's routine working: and paused: appends woke the primary every time. signal_crew_provably_working now reads the mate's lines new since the watcher's classified position: a decision, blocker, terminal outcome, note:, correlation-marked line, or unknown verb still surfaces regardless of busy evidence, while unmarked working:, paused:, and resolved: fall through to the same provably-working absorb an ordinary crewmate gets. * no-mistakes(review): narrow secondmate routine absorb to working and paused
* fix(bin): use gh-axi for the ship DoD draft check * no-mistakes(review): use PR number not URL in gh-axi draft check
… only (kunchenguid#5520) * fix(bin): drop status prose from the inactive-outcome dedupe identity The inactive-outcome receipt fingerprint included the child's sanitized last status line, so a persistent child appending routine prose after one terminal outcome minted a fresh parent event per sentence. Bind the identity to incarnation, task id, terminal state, and PR only, keeping the last line in the record as status_head evidence. Fixes kunchenguid#2960 * no-mistakes(document): note structured-only inactive receipt identity in regression coverage
…unchenguid#5707) * feat(bin): record the supervision host's dialog mirror on Claude and Cursor Add bin/fm-host-mirror.sh, the one owner of the supervision host's dialog mirror file, cursor, lock, and feed, plus the main-session key it keys entries to. The tracked Claude UserPromptSubmit and Stop hooks and the Cursor beforeSubmitPrompt and afterAgentResponse hooks record the captain's prompt and main's reply, only on a home with config/supervision-host, from a genuine primary checkout, for the lock-owning session. The mirror lands inert: writers record and nothing reads it yet; attended supervision on the host is the later step that consumes the feed. Codex, Grok, OpenCode, and omp have no writer here. * no-mistakes(review): Scope mirror dedup to session, atomic appends, marker-inclusive caps * no-mistakes(document): Clarify dialog mirror scope and remove duplicate contract details * no-mistakes(document): Correct Cursor hook documentation for dialog mirror registration * no-mistakes(review): Pass mirrored dialog text to jq via stdin * no-mistakes(document): Clarify dialog mirror documentation and remove duplicate claims * no-mistakes(review): Preserve internal dialog whitespace; drop mirror check and verified modes * no-mistakes(review): Drop only identical mirror repeats; remove redundant chmod guard
… escalations are not repeated (kunchenguid#5731) * fix(bin): retire check-row receipts on branch acks and report an unchanged situation once * fix(bin): scope a branch acknowledgement's check-row receipt retirement to its granted sequences The away posture lifts the attended partition's check/decision exclusions, so a branch grant can name check-kind rows - but the branch-actor ack still assumed check rows were main-only and skipped every receipt scan. The queue row was consumed while its terminal-outcome .pending receipt stayed behind, and each inactive-reconcile cadence scan re-queued the same fingerprint. In the first real away window on the supervision host that re-escalated one unchanged held-PR situation on every cycle (~1,734 of 4,149 outcomes). A branch ack now scans inactive-outcome and inactive-reconcile receipts and commits secondmate stall receipts against exactly the sequences in its eligible-row snapshot - the same rows it consumes - instead of none. Attended grants still name no check row, so the scans find nothing. * fix(bin): store a repeated captain verdict as routine while the task's durable situation is provably unchanged fm-branch-outcome.sh append computes a mechanical situation key per captain row - metadata bytes, captured status-log endpoint and identity, live crew-state verb, worktree head - and anchors it in state/.<task>.branch-captain-key. A later captain verdict whose recomputed key matches is stored as routine with "unchanged since seq <N>:" prefixed to its summary, so one situation escalates once until something provably changes. A task with no readable status ledger is never demoted, an unreadable record fails toward reporting, and teardown removes the sidecar with the task's other branch records. The append-only store schema is unchanged. This covers both hosts: the Pi supervision branch and the supervision host both funnel reports through append. * docs: check rows are main-owned only while attended; the away posture grants them to the branch, whose ack retires their receipts exactly * test: the away-flood reproduction as a regression test (branch ack retires the receipt and later scans stay quiet), store-level dedupe coverage, and a branch-ack secondmate stall receipt case * fix(bin): restore the secondmate child devin-config cleanup path The branch-captain-key sidecar addition mistyped the sibling entry as .$child_id.devin-config.json, so a forced secondmate teardown would have stopped removing each child's real <id>.devin-config.json. Restore the original path and add a behavioral test that stops the child sweep mid-loop on a refused close, proving the cleaned child's devin config and captain anchor are both removed while the unconsumed child's records are retained. * no-mistakes(review): Key captain dedupe on the covered wake rows' fingerprint * no-mistakes(review): Drop captain-key demotion; prove one escalation on both surfaces * no-mistakes(review): Drop unrelated teardown test; cite both receipt test files * no-mistakes(document): Docs already match branch-ack check-receipt retirement
…oorbell (kunchenguid#5664) * fix(calm): deliver Claude-bound operational input as a record-backed doorbell Claude Code 2.1.280 removes U+2063 from every submitted prompt, so a typed operational envelope reaches a Claude Code primary as plain text. The away daemon now writes the envelope to a record under state/operational-inbox and types only a plain doorbell naming it; the /afk return check and the Calm mod recognize the doorbell only when that record holds a current envelope. Marker- preserving harnesses keep the typed envelope. The live Calm guard accepts the 2.1.280 module-load log line, drives the doorbell, and asserts thinking stays hidden. * no-mistakes(review): Fix operational record retention at 7 days and document prune limit * no-mistakes(document): Point Calm bounds at 2.1.280 evidence; fix afk-exit comment * no-mistakes(lint): Pick newest Calm e2e transcript without parsing ls * docs(calm): add a minimal turning-Calm-on step for Claude Code * fix(spawn): deliver the Claude launch brief as a record-backed doorbell Claude Code strips U+2063 from the launch-prompt argument too, so a worker's launch brief arrived with its operational marker removed. Publish the brief as a record in the receiving home's operational inbox - a secondmate's own state, not the primary's - and pass only the printable doorbell naming it, falling back to the typed envelope when the record cannot be published so the brief body still delivers. Unwrap doorbell-carried digests in the daemon digest tests that still read the raw send log under the claude pin, and update the documented bounds now that launch briefs hide like the other operational rows. * test(spawn): cover a secondmate's launch-brief record landing in its own home The record-backed doorbell resolves its state through the receiving pane's home, so prove a claude secondmate launch publishes into the seeded secondmate's operational inbox and never leaks a record into the primary's. * no-mistakes(review): Pass primary harness to daemon, tighten retention, refresh verdicts * no-mistakes(review): Prune operational records by exact seven-day elapsed age * no-mistakes(review): Batch record pruning so large inboxes still expire * no-mistakes(review): Refuse Claude spawn when brief record cannot publish * no-mistakes(review): Drop thinking probe from Claude Calm live test and docs * no-mistakes(review): Record dated Claude Code 2.1.282 reproduction evidence * no-mistakes(document): Clarify operational doorbell documentation and record expiry * no-mistakes(document): Correct AFK escalation carrier guidance * no-mistakes(review): Describe operational record retention as about seven days * no-mistakes(document): Clarify Calm delivery and operational record retention * no-mistakes(review): Remove out-of-scope Calm launch guide from Claude docs * no-mistakes(document): Document Claude launch-brief delivery and refusal * no-mistakes(document): Correct stale operational-input documentation * no-mistakes(ci): Fixed the stale Claude trust test to verify that worker and secondmate launches deliver readable, record-backed briefs instead of expecting brief paths in their commands. Annotated the daemon’s output variable for ShellCheck without changing behavior. The affected tests, daemon tests, ShellCheck, and diff check pass locally * no-mistakes(ci): parse rebased Claude launch after trailer hook prefix * no-mistakes(review): Trust launch-brief record and restore thinking bound doc * no-mistakes(review): Parse final Claude launch statement; drop Stop-hook docs --------- Co-authored-by: Mike Sewell <maikunari@protonmail.com> Co-authored-by: no-mistakes <no-mistakes@localhost>
kunchenguid#5583) A host-local relaunch rewrote only the far endpoint, so this home kept the old harness, model, and effort, and appending those keys after pr= broke pull-request poll authentication.
…ound (kunchenguid#5516) tests/fm-watch-triage.test.sh finishes in about 434s alone and about 698s under CI load, so the 900s bound the changed-suite runner applies produced a false timeout under ordinary concurrent validation. Raise the automatic bound to 1500s, which keeps every measured script under it while staying below the 30-minute normal CI tier so a genuinely hung script still fails here with its output before the job cap cancels the lane. Fixes kunchenguid#3869 Refs kunchenguid#3565
…kunchenguid#5728) * Fix nested watcher lock reclaim * no-mistakes(review): Elect a single steal-mutex reaper and bound arm TERM wait * no-mistakes(review): Reclaim self-held steal mutex and unify autoarm steal reaping * no-mistakes(review): Resume own interrupted steal reap from its tombstone
…nguid#5710) * test: hold the back-to-back boundary close on the host's own clock test_park_boundary_holds_under_back_to_back_closes assumed two engine turns fit in the ~16s pre-refusal window and that the stub finished a turn in 3s. Under load the stub's real drain, report, and acknowledgement take ~13s, so the turn either died at its bound (which hands the wake to main, no boundary line) or the second close landed past the window and the fixture failed while the boundary held. 3 failures in 5 runs at a load average near 11. Hold the first turn on a release file instead: once the engine is in flight, a second close is appended mid-turn and the turn is released as the refusal window opens (park bound minus turn bound and grace, read off the host's own start record). The queued close can then only wait for the boundary on any machine speed, which is what the test asserts: the boundary line ends the output, the demo.status row stays queued for main, and no second engine turn ever starts. A host too loaded to start the turn at all hands the first close to the same boundary exit. After: 12/12 at load ~15-42. * no-mistakes(review): Print boundary test deadline as a decimal integer * no-mistakes(review): Hold boundary test turn on a FIFO, require full sequence * no-mistakes(review): Remove stray before/after supervision-host test copies * test: hold the late close's render until the refusal window opens The boundary recheck test's node shim slept a fixed 10s, which assumed the first close was read before the host's refusal window opened. Under load the close arrived after the refusal check, so the host correctly refused it before the successor started and the render snapshot never appeared. Block the wake-prompt render on a FIFO released at the refusal-open instant read from the host's own start record, so the pre-turn recheck must refuse on any machine speed. * no-mistakes(review): Derive minimal park bounds and refresh supervision-host shard hint * no-mistakes(review): Drive park-boundary tests from a seam-gated host test clock
* fix(bin): stage remote home clones before publishing them A remote home provision cloned the code root directly into the public FM_HOME path while rollback() claimed rm -rf of that same path on any failure. Bash defers trapped signals past a foreground child, but any other cleanup or lifecycle path that removes the home directory races the live clone's object copy, producing the CI flake "fatal: failed to copy file to .../.git/objects/...: No such file or directory". Clone into a private staging directory beside the home and publish with an atomic rename once complete, so no cleanup can remove a directory a live clone is still writing; a home that appears mid-provision now dies cleanly instead of inheriting torn state. The regression coverage holds a real clone mid-copy, removes the public path, and requires the provision to finish and publish intact. * no-mistakes(review): Prove home ownership by sentinel and hold only a live clone * no-mistakes(review): Assert raced provision publishes a complete, intact clone * no-mistakes(document): Document remote home staging and publication safety * no-mistakes(lint): Fix ShellCheck warning in clone integrity assertion * no-mistakes(document): Clarify remote home publication and rollback guarantees
* fix(control): keep a relaunched Pi worker's herdr pane status authority alive Defect: after `bin/fm-control.sh <id> relaunch` (observed live on a herdr Pi crewmate whose pane read idle while it ran its validation pipeline), the pane froze at whatever its previous agent had last reported. Cause, measured on herdr 0.9.1 against a real Pi: a pane has one status authority, and for Pi with its integration installed that authority is the lifecycle hooks, so herdr also skips screen detection for the pane. In the crew shape the registration outlives its agent process (upstream issue kunchenguid#4115; docs/herdr-backend.md "Restart and liveness behavior"), and herdr applies only reports carrying the session identity it bound. A replacement started fresh in that pane reports a NEW session, so its state reports are ignored and the pane stays frozen. Nothing from outside repairs it: `pane report-agent-session` and `pane report-agent` for `herdr:pi` are accepted (rc=0) without being applied unless the reporter is the registered pane agent, and `pane release-agent` on the stale record changes nothing. Fix: a relaunch preserves the binding instead of fighting it. The launch owner reads the session reference the endpoint's own runtime recorded (`fm_backend_herdr_pane_agent_session_ref`) and passes it back as Pi's own `--session <path-or-id>` (`relaunch_resume_args`; `fm_control_relaunch_resume_flag` owns which adapters and which registered-agent labels qualify). That is the same reference herdr itself resumes Pi panes with after a server restart, and the resumed session's reports land again, which the live check confirmed: the pane returned to working while the replacement worked and idle when it settled, on the same session identity. Safety: relaunch-only (a fresh spawn binds nothing), herdr-only (the one adapter that records a per-pane session), Pi-family only, and only when the registration's own agent label matches - so no other adapter's conversation can be handed to a Pi launch. An unreadable, missing, or malformed reference degrades to exactly the fresh-session launch that existed before. No lifecycle, liveness, isolation, or merge guard is touched, and an empty result leaves every non-Pi launch byte-identical. `resume` remains a refused verb; docs/agent-control.md and the harness-adapters references are corrected where they claimed Pi had no verified resume form at all. * no-mistakes(document): docs: correct relaunch session-authority ownership and skill paths * no-mistakes(document): docs: correct stale control-plane ownership claim * no-mistakes(document): docs: drop unverified Herdr restart resume claim * no-mistakes(test): Added offline Herdr Pi session-authority relaunch coverage * no-mistakes(document): Document Herdr Pi relaunch session continuity * no-mistakes(ci): The failing remote relaunch test tried to arm a PR poll for a secondmate, which `fm-pr-check.sh` correctly refuses. Removed that invalid test scenario; the remaining remote relaunch tests pass, and `git diff --check` is clean
…nchenguid#5758) Main has been red since fm-pr-check.sh began refusing to arm a merge poll on a kind=secondmate record (kunchenguid#5696): the relaunch-ordering case in tests/fm-remote-secondmate-relaunch.test.sh armed its fixture through that entry point and could no longer be set up. The ordering guarantee still matters: a secondmate record armed before the refusal can legitimately carry a trailing pr=/pr_head= identity block until the watcher retires it, and fm-remote-secondmate-relaunch.sh must still keep that block last when republishing harness/model/effort. Seed the fixture the way such a record was really written - pr= appended last to the meta, then the poll artifacts published through the same fm_pr_poll_prepare/fm_pr_poll_publish_prepared pair fm-pr-check.sh uses, a pattern tests/fm-pr-check-security.test.sh already follows - and drop the now-unused fake gh fixture. The kunchenguid#5696 refusal itself stays pinned by the security suite's secondmate-record case.
…id#5748) * feat: run attended supervision on the host for Claude and Cursor On a home opted into config/supervision-host with a Claude or Cursor primary, the supervision host now takes the attended wakes the Pi branch would take: routine outcomes stay off main, and a captain outcome wakes main once with a branch-outcome line and waits in the drain's new BRANCH OUTCOMES section until main acknowledges it with mark-processed. - The offer rule moves into branchOfferForWake, shared by the Pi watcher and the host through bin/fm-branch-dispatch.mjs offer. - The host feeds the dialog mirror at the head of each attended wake and passes a close through unchanged when it is main-only, the engine or a tool is missing, the primary has no verified mirror, the main session cannot be identified, or the session is cooling down. - The drain presents captain outcomes first, one line per task, never behind older routine outcomes, and collapses routine overflow into a count that is marked read. - The return advances the store's read cursor through the away window once the brief has rendered, so the first drain does not replay it. - The branch prompt's mirror wording is host-neutral, and the rule to report what main must act on as captain, once per unchanged situation, applies only to the attended posture on the host. * docs: record the attended supervision host live check * no-mistakes(review): Present pre-window unread outcomes and contiguous captain prefix * no-mistakes(review): Return brief presents every row it marks read * no-mistakes(review): Return brief lists every unread outcome in one list * no-mistakes(review): Keep return list in store order and gate cursor failures * no-mistakes(review): Make the drain the only branch-outcome presenter after return * no-mistakes(review): Gate return on drain outcome failures; byte-count outcome budgets * no-mistakes(review): Gate drain on projection failures; UTF-8-safe byte cuts * no-mistakes(review): Fail drain without jq; hand unreadable prompt mirror to main * no-mistakes(document): Correct supervision-host return and drain documentation * no-mistakes(review): Recheck attended offer at turn start; honest failed-drain brief * no-mistakes(document): Correct supervision-host posture and drain documentation * no-mistakes(document): Documentation remains accurate for attended supervision
…unchenguid#5753) Each '# shellcheck source=' directive makes ShellCheck's external-source traversal expand that library's whole transitive graph again at the site. fm-pending-reply-lib carried three directed lazy sources of fm-wake-lib and two of fm-parent-channel-lib on identical per-call re-source sites, so one file analysis peaked above 4 GiB and every caller (fm-watch, fm-teardown) inherited the multiplier - the root cause of the PR kunchenguid#5732 Lint 1 OOM kill. Keep the runtime '.' commands byte-identical: the lazy re-source under 'local STATE FM_WAKE_QUEUE FM_WAKE_QUEUE_LOCK' is real behavior. Drop the duplicate directives so each library expands once per unit, and drop the tmux/classify directives since classify already arrives through the kept fm-wake-lib expansion and no tmux symbol is referenced here. The directive above the lib-dir assignment is kept - it binds the bin/ prefix so the undirected sites still resolve without SC1091. Measured peak RSS, ShellCheck 0.11.0 -x on Linux arm64: bin/fm-pending-reply-lib.sh 4.06 GiB -> 1.96 GiB, zero findings
kunchenguid#5773) * fix: split bash 5.2 sibling $() in recovery mint and delivery log Sibling command substitutions on one line can empty a recovery generation under bash 5.2 when a CHLD trap is set. Mint pid/epoch sequentially, refuse empty tokens before write, and clean delivery fields before printf. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: split bash 5.2 sibling $() in recovery mint and delivery log Sibling command substitutions on one line can empty a recovery generation under bash 5.2 when a CHLD trap is set. Mint pid/epoch sequentially, refuse empty tokens before write, and clean delivery fields before printf. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: keep recovery mint failure semantics after sibling $() split Remove the new pid/date refusal and grammar guard so a mint miss still yields a grammar-valid token and a durable wake row, matching accepted review intent. Drop the fake-failing-date case that locked in the refuse. Co-authored-by: Cursor <cursoragent@cursor.com> * no-mistakes(document): Point recovery-mint hazard comment at its regression test --------- Co-authored-by: Cursor <cursoragent@cursor.com>
…stop calling Escape safe there (kunchenguid#5791) Fixes kunchenguid#4520
…unchenguid#5790) Fixes kunchenguid#4756 The voice status reader in bin/fm_voice_records.py reports each worker's state from the last non-blank line of its status log. When a worker appends a status line and then a line of plain prose, the reader reported "note" with the prose line instead of the declared state, diverging from bin/fm-classify-lib.sh's shell scan. Scan back through the tail for the newest line whose prefix is a single lowercase verb-shaped word (letters and hyphens), and report that event's verb instead of always taking the last line. An unrecognised verb-shaped prefix still reports "note" rather than letting an earlier recognised line answer for it, and free text with no colon is skipped as prose. When the tail holds no such event, the last line is reported exactly as before.
…newer version (kunchenguid#5786) * fix(bin): stop reporting an already-installed version as an available update An update announcement named its version first ("current -> new"), so reading the first dotted number as the announced version compared the current version against itself and always looked newer. Read the last dotted number instead, and only report an available update when that announced version is newer than the newest installed copy found; when that version is already installed, report only PATH skew. Fixes kunchenguid#5151 * no-mistakes(document): docs: gate announce update-available report on newer-than-installed
Conflicts resolved preferring upstream's mechanism where it covers the fork's case; see the PR body for the per-file table.
--check --threshold N prints one line only when the fork is more than N upstream commits behind, once per drift episode, recording the episode in state/.upstream-drift because the watcher does not deduplicate check output. CONTRIBUTING documents the weekly merge cadence and how to arm the check.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Merges
upstream/main(kunchenguid/firstmate, 173 commits throughe9a6675e, 2026-09-26) into the fork, the first sync since 2026-09-13, and arms a drift alert so the fork no longer falls silently behind.Land this with a MERGE commit (
fm-pr-merge.sh -- --merge), never squash or rebase. Every home fast-forwards; a squash would drop upstream's history and make the next sync re-conflict on every hunk.What changed
Merge upstream/main into fork- a real two-parent merge made withgit rerereon (-c rerere.enabled=true -c rerere.autoupdate=true, the same flagsbin/fm-upstream-sync.shuses; run in this task's isolated worktree rather than the script's own$TMPDIRworktree). 15 files conflicted. Rule applied: where both sides fixed the same failure, upstream's version wins and the fork's is dropped; fork code is kept only where upstream does not cover the fork's case.feat(bin): add an upstream drift alert to fm-upstream-sync --check---check --threshold Nprints one line only when the fork is more than N upstream commits behind, once per drift episode. CONTRIBUTING gains a "Syncing a fork" section with the weekly cadence and the arming steps.Conflict resolution, per file
bin/fm-session-lock-lib.shdaemon run/bg-sparelayers in the ancestry walk with aCLAUDE_PIDexception; strictfm_harness_pid_aliveexcluding those layers; own.lock-sessionidentity (fm_session_lock_owned_by_session,fm_session_lock_owner_alive,fm_harness_pid_live_any)CLAUDE_CODE_SESSION_IDcounted only whenCLAUDE_PIDis a Claude-shaped member of the caller's contiguous ancestry),fm_session_lock_anchor_pid(=CLAUDE_PID), no layer skipping,fm_session_lock_inspect,fm_session_lock_foreign_owner_liveCLAUDE_PID; its anchor is the model-loop pid, never the daemon. Proof below.bin/fm-lock.shbin/fm-claude-stop-autoarm.shfm_session_lock_pid_is_self/owner_alive; #15 need gate usesfm_supervision_arm_neededbin/fm-turnend-guard-cursor.sh(clean merge, fixed up)fm_session_lock_owner_alive; #15fm_supervision_arm_neededfm_harness_pid_alive; #15 kept.bin/fm-wake-lib.shFM_LOCK_STEAL_TIMEOUT,FM_LOCK_STEAL_RETRY_DELAY, "lock steal contention exhausted")fm_lock_try_acquire_steal_mutex+fm_lock_reap_dead_link(tombstone-elected reaper, no nested steal mutex)bin/fm-pr-merge.shbranches/<b>/protection, plan-limit 403 as "none required",--allow-missing--allow-missingbin/fm-send.sh--resolve-keycombines with--fire-and-forget(closes an escalated pending-reply decision)fm-send-resolve-keycovers it.tests/fm-session-lock-ancestry.test.shfm_session_lock_anchor_pid,fm_session_lock_owned_by_self,fm_session_lock_trusted_session_id) instead of fork-internal helpers; fork fixture knobs (FM_FIXTURE_*_ARGS,FM_FIXTURE_DECLARE_SESSION) re-added.CLAUDE_PID), which upstream no longer has, and #5's "legacy daemon-pid lock reads dead" assertion (no lock can record a daemon pid under upstream's anchor).tests/fm-watcher-lock.test.sh(clean merge, fixed up)test_lock_steal_contention_is_boundedkept minus the fork-only message/env.tests/fm-pr-merge.test.shtest_github_required_checks_refuse_missing_allow_named_and_fail_closedtests/fm-captain-hold-lifecycle.test.sh/protectionendpointdocs/architecture.mdsilent_idle_check) is fork-only and still infm-watch.sh.docs/configuration.mddata/<id>/timeline.json.docs/scripts.mdfm-upstream-sync.shrowfm-update.shrow rewordeddocs/turnend-guard.md,docs/watcher-continuity.md,docs/verification/supervision.mdOther audit overlaps that merged cleanly: fork #14 (stop re-alarming a delivered PR each turn end) is kept - upstream kunchenguid#5599 fixes a different failure (repeated unknown-wake escalations in the away daemon). Fork #11's silent-idle backstop and
--resolve-keyremain.Fork-only features intact:
fm-host-report.sh,fm-mate-view.sh,fm-bridge-snapshot.sh/fm_bridge_snapshot.py,fm-bridge-console.py,fm-task-timeline.sh/fm_task_timeline.py,fm-context-check.sh,fm-upstream-sync.sh.git diff upstream/main HEADis now 49 files, all fork features or the kept fork fixes above.Proof: upstream kunchenguid#4894 covers fork #13 (background-hosted session)
tests/fm-session-lock-ancestry.test.shagainst upstream's lock lib, fork cases:The last one is fork #13's real-process regression (orphaned
bg-pty-hostpassing--bg-sparethrough, hosting the claimed spare that is the session, no interactive claude above): the real Stop auto-arm claims the home, arms, and the lock stays on the session pid.Drift alert
bin/fm-upstream-sync.sh --check --threshold 50printsupstream: N new commits not in origin/main (fork sync due; drift alert past 50)once when the fork falls more than 50 commits behind, then stays silent (recorded instate/.upstream-drift) until a sync brings it back within 50. The watcher does not deduplicate check output and sweeps every 300 s, so a plain--checkwould re-wake every sweep.config/watched-tools.jsonwas not used: its git probe compares a clone with its own remote and never fetches, so it cannot express fork-vs-upstream drift with a threshold.Arming is a post-merge step in the primary home (this worker cannot write outside its worktree, and the home's
bin/only gains--thresholdafter merge): write mode-0700state/upstream-drift.check.shexportingFM_HOMEandFM_ROOT_OVERRIDEas the home and exec'ingbin/fm-upstream-sync.sh --check --threshold 50, thenbin/fm-check-register.sh upstream-drift. CONTRIBUTING "Syncing a fork" documents it and the weekly cadence.Validation
tests/fm-session-lock-ancestry.test.sh,tests/fm-watcher-lock.test.sh(42 ok),tests/fm-upstream-sync.test.sh;bin/fm-lint.shclean on every changed bin script and test;bin/fm-test-run.sh --check-coverageok (241 tests);bin/fm-doc-audience-check.shok.bin/fm-test-run.sh --all) was stopped after 35 of 241 scripts; 19 of those 35 failed, all in herdr/backend-smoke, afk, and backlog/bearings families that drive a live terminal multiplexer or home state on this host (the run was inside a live herdr session). Not triaged against a stock-upstream baseline; fork CI is the authority.Follow-ups (not in this PR)
AGENTS.mdgrows from 80,569 B to 90,081 B (+9,512 B, ~2.4k tokens per call, 12,236 words against upstream's own 9,000-word ceiling), almost all upstream's. Pruning is deliberately out of scope here.